Papers with social sciences

16 papers
Measuring and Modeling Language Change (N19-5)

Copied to clipboard

Challenge: This tutorial will help researchers answer questions fundamental to the social sciences and humanities .
Approach: This tutorial is designed to help researchers answer questions in the social sciences and humanities . it synthesizes recent computational techniques for handling and modeling temporal data .
Outcome: The tutorial will synthesize recent techniques for handling and modeling temporal data, such as dynamic word embeddings, and identify useful tools for social scientists and digital humanities scholars.
Do Large Language Models Discriminate in Hiring Decisions on the Basis of Race, Ethnicity, and Gender? (2024.acl-short)

Copied to clipboard

Challenge: We study whether large language models exhibit race- and gender-based name discrimination in hiring decisions .
Approach: They propose templatic prompts to LLMs to write an email to a named job applicant informing them of a hiring decision.
Outcome: The proposed model generates an acceptance or rejection email based on the applicant's first name .
A Query-Driven Topic Model (2021.findings-acl)

Copied to clipboard

Challenge: Topic modeling is an unsupervised method for revealing the hidden semantic structure of a corpus.
Approach: They propose a query-driven topic model that allows users to specify a simple query in words or phrases and return query-related topics.
Outcome: The proposed model is particularly attractive when the query has a low occurrence in a text corpus, making it difficult for traditional topic models to identify relevant topics.
ILCM - A Virtual Research Infrastructure for Large-Scale Qualitative Data (L18-1)

Copied to clipboard

Challenge: iLCM project develops integrated research environment for qualitative data analysis . text mining and text mining tools are extended by "Open Research Computing"
Approach: iLCM project develops integrated research environment for analysis of structured and unstructured data in a "Software as a Service" architecture.
Outcome: iLCM project develops integrated research environment for analysis of structured and unstructured data in a "Software as a Service" architecture.
Can Large Language Models Discern Evidence for Scientific Hypotheses? Case Studies in the Social Sciences (2024.lrec-main)

Copied to clipboard

Challenge: scholarly databases fail to aggregate, compare, contrast, and contextualize existing studies in service to a targeted research question.
Approach: They propose to use large language models to discern evidence in support or refute of specific hypotheses based on abstracts.
Outcome: The proposed method outperforms state-of-the-art methods and highlights opportunities for future research.
Ideology Takes Multiple Looks: A High-Quality Dataset for Multifaceted Ideology Detection (2023.emnlp-main)

Copied to clipboard

Challenge: Existing datasets for the ID task only label a text as ideologically left- or right-leaning as a whole, regardless whether the text containing one or more different issues.
Approach: They construct an ideological schema for a multifaceted ideology detection task using MITweet and an English Twitter dataset.
Outcome: The proposed task uses a MITweet dataset with 12,594 English Twitter posts, each annotated with a Relevance and an Ideology label for all twelve facets.
Hong Kong: Longitudinal and Synchronic Characterisations of Protest News between 1998 and 2020 (2022.lrec-1)

Copied to clipboard

Challenge: This paper examines the utility and timeliness of the Hong Kong Protest News Dataset . it sheds light on whether depth and/or manner of reporting changed over time .
Approach: They use the Hong Kong Protest News Dataset to investigate synchronic news characterisations of protests in Hong Kong between 1998 and 2020.
Outcome: The dataset sheds light on whether depth and/or manner of reporting changed over time, and if so, in what ways, or in response to what.
Conceptualizing Treatment Leakage in Text-based Causal Inference (2022.naacl-main)

Copied to clipboard

Challenge: Existing methods to control for text-based confounders rely on assumption that there is no treatment leakage . prior literature has assumed that documents only contain information about confounder, but not about treatment assignment.
Approach: They define the treatment leakage problem and propose methods to mitigate it . they remove treatment-related signal from text in a pre-processing step .
Outcome: The proposed method can mitigate the problem of treatment leakage by removing the treatment-related signal from the text.
Perhaps PTLMs Should Go to School – A Task to Assess Open Book and Closed Book QA (2021.emnlp-main)

Copied to clipboard

Challenge: Taking the exam closed book, but having read the textbook, yields at best minor improvement (56%), suggesting that the PTLM may not have “understood” the textbook (or perhaps misundersttoo the questions).
Approach: They propose to use pre-trained language models to answer questions from introductory college textbooks and hundreds of true/false statements based on review questions written by the authors.
Outcome: The proposed task includes two college-level introductory texts in the social sciences (American Government 2e) and humanities (U.S. History).
The Role of Pragmatic and Discourse Context in Determining Argument Impact (D19-1)

Copied to clipboard

Challenge: Recent work shows that attributes of both the audience and communicator constitute important cues for determining argument strength.
Approach: They propose to use a dataset to study the pragmatic and discourse context of argumentative claims to build predictive models that incorporate the pragmatic context of the argument.
Outcome: The proposed models outperform models that rely on claim-specific linguistic features for predicting the perceived impact of individual claims within a particular line of argument.
BnMMLU: Measuring Massive Multitask Language Understanding in Bengali (2026.findings-acl)

Copied to clipboard

Challenge: Large-scale multitask benchmarks have driven rapid progress in language modeling, yet most emphasize low-resource languages like English.
Approach: They propose a benchmark for massive multitask language understanding in Bengali . they use a dataset that preserves mathematical content via MathML and a subset of questions most frequently missed by top systems to stress difficult cases.
Outcome: The proposed benchmark covers 24 model variants across 11 LLM families.
Discovering influential text using convolutional neural networks (2024.findings-acl)

Copied to clipboard

Challenge: Existing methods for estimating the effects of text on human evaluation are limited to testing a small number of pre-specified text treatments.
Approach: They propose a method for flexibly discovering clusters of similar text phrases that are predictive of human reactions to texts using convolutional neural networks.
Outcome: The proposed method can detect and predict human reactions to texts under certain assumptions.
Tab2Text - A framework for deep learning with tabular data (2024.findings-emnlp)

Copied to clipboard

Challenge: Tabular data is a foundational part of social sciences and is used to fit supervised learning models.
Approach: They propose a technique for transforming tabular data to text data to improve deep learning models for tabular datasets.
Outcome: The proposed technique improves performance of deep learning models for tabular data.
From Script to Stage: Automating Experimental Design for Social Simulations with LLMs (2026.findings-acl)

Copied to clipboard

Challenge: Xu et al., 2024): multi-agent simulations based on large language models are a new paradigm for social science research . traditional experimental design relies on interdisciplinary expertise and technical barriers . Xiaoping and Xin eli argue that LLM-driven agents are unreliable for rigorous experimental design due to hallucinations and limited verifiability.
Approach: They propose a framework for multi-agent experiment design based on script generation . Script Composition, Script Finalization, and Actor Generation are the core phases of the framework .
Outcome: The proposed framework lowers the barrier for social science experimental design and provides scientifically grounded decision support for policy-making.
PSE v1.0: The First Open Access Corpus of Public Service Encounters (2024.lrec-main)

Copied to clipboard

Challenge: a dataset of public service encounters in germany provides a new research directive . data from the public service encounters are used to investigate bias, bureaucratic discrimination and other power-driven dynamics in the actual communication .
Approach: They propose to compile a dataset of transcribed public service encounters in germany . they propose to open up the black box of direct state-citizen interaction .
Outcome: The proposed dataset allows the community to open up the black box of direct state-citizen interaction.
The ParlaSent Multilingual Training Dataset for Sentiment Identification in Parliamentary Proceedings (2024.lrec-main)

Copied to clipboard

Challenge: The paper presents a new training dataset of sentences in 7 languages, manually annotated for sentiment, which is used in a series of experiments focused on training a robust sentiment identifier for parliamentary proceedings.
Approach: They propose to use a dataset of sentences manually annotated for sentiment to train a robust sentiment identifier for parliamentary proceedings.
Outcome: The proposed model performs very well on languages not seen during fine-tuning and additional fine- tuning data from other languages significantly improves the target parliament’s results.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations